Phase 3: Machine Learning Lesson 3 of 6

Supervised Learning:
Regression

Not every prediction needs to be a category. Sometimes you need a number: a price, a temperature, a score. Regression is supervised learning for continuous outputs, and it is one of the oldest and most reliable tools in the entire field.

You will learn
The difference between regression and classification
How linear regression works geometrically
What loss functions measure in regression
MAE, MSE, and R-squared explained
Predicting house prices in Scikit-learn

Regression versus classification

The first question to ask when you start a machine learning project is: what kind of thing am I trying to predict? If the answer is a label or category, you reach for a classifier. If the answer is a number somewhere on a continuous scale, you reach for a regressor.

Regression predicts...
  • House price ($312,000)
  • Temperature tomorrow (23.4 °C)
  • Patient's blood pressure (128 mmHg)
  • Time a delivery will take (47 minutes)
  • Number of sales next quarter (8,420 units)
Classification predicts...
  • Will this loan default? (Yes / No)
  • What species is this? (Cat / Dog / Bird)
  • Is this email spam? (Spam / Not spam)
  • What digit is this? (0 through 9)
  • Is this tumour benign? (Benign / Malignant)

The boundary can get fuzzy. You can turn a regression problem into a classification one by bucketing the output: instead of predicting the exact house price, predict whether it is "cheap," "mid-range," or "expensive." But for anything where the exact number matters, regression is the right tool.

Linear regression: fit a line, predict a number

Linear regression is the simplest regression algorithm, and it is genuinely useful. The core idea is that you draw a straight line through your data in a way that minimises the total distance between the line and all your data points. Future predictions are just reading off the line at the relevant x-value.

Analogy

You are studying for an exam and you have records from previous students: how many hours each person studied and what score they got. You plot this as a scatter of dots. Linear regression draws the single best straight line through that scatter. Once the line is drawn, you can read off "if I study 7 hours, I should expect a score of about 74."

The equation of a line (linear regression)
y = mx + b
y = the prediction (what we want) x = the input feature m = slope (how steep the line is) b = intercept (where the line crosses y-axis)

The model's job during training is to find the values of m and b that make the line fit the training data as well as possible. These are the parameters of the model, the numbers the model learns. Once training is done, m and b are fixed. Predicting is then just plugging in any x and computing the result.

When you have multiple input features, the line becomes a hyperplane, but the principle is identical. Instead of one slope m, you have one weight (slope) for each feature. Still the same core idea: find the weights that minimise prediction error.

Linear regression: study hours vs exam score
Study hours Score 0 2 4 6 8 10 Best fit line Residuals (errors)

The green line is the best fit through the data. The orange dashed lines are residuals: the vertical distance between each point and the line. Linear regression finds the line that minimises the total size of these residuals.

Loss functions: measuring how wrong you are

To fit the line, the model needs a way to measure how wrong its current parameters are. That measure is the loss function. In regression, the most common approach is to look at the residuals: the gap between each prediction and the actual value. The loss function summarises all those gaps into a single number.

Three metrics dominate regression evaluation. Understanding what each one rewards and what each one punishes will save you from misinterpreting your model's performance.

MAE
Mean Absolute Error
Average of the absolute differences between predictions and actuals. Easy to interpret: an MAE of 5,000 means your predictions are wrong by $5,000 on average. Treats all errors equally.
MSE
Mean Squared Error
Average of the squared differences. Squaring penalises large errors much more harshly than small ones. Used during training because it is mathematically smooth to optimise. Harder to interpret in original units.
R²
R-Squared
Measures how much of the variance in the target the model explains. Ranges from 0 to 1 (higher is better). An R² of 0.85 means the model explains 85% of the variability in house prices.
Which metric to report?

Use MAE when you want an error in the same units as the target (dollars, degrees, minutes) and need to explain the result to a non-technical audience. Use MSE or RMSE (root of MSE) during model training and comparison. Report R-squared alongside MAE when presenting to stakeholders: "The model explains 85% of price variation and is off by an average of $12,000."

When a straight line is not enough

Linear regression assumes the relationship between input and output is a straight line. Real-world relationships often curve. A student's scores might improve steadily with study hours up to a point, then plateau due to fatigue. A car's fuel efficiency might drop slowly at moderate speeds then plummet at highway speeds.

Polynomial regression handles this by adding powers of the feature as extra columns. Instead of just using x, you also include x squared, x cubed, and so on. The model is still linear in its parameters (it learns weights for each term) but the resulting curve can bend and flex to fit non-linear patterns.

More complex relationships call for more powerful algorithms. Decision trees can do regression (predict numbers at each leaf instead of class labels). Random Forests average the predictions of many trees. Gradient Boosting builds trees sequentially, each one correcting the errors of the last. These ensemble methods consistently rank among the best performers on tabular data with complex, non-linear relationships.

Predicting house prices in Scikit-learn

The classic regression toy problem is predicting house prices from features like size, location, number of rooms, and age. Here is a full pipeline using the California housing dataset that ships with Scikit-learn.

Python house_price_regression.py
from sklearn.datasets import fetch_california_housing
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error, r2_score
import numpy as np

# Load the built-in California housing dataset
data = fetch_california_housing()
X, y = data.data, data.target

# Split into training and test sets
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train a linear regression model
model = LinearRegression()
model.fit(X_train, y_train)

# Predict on the test set
predictions = model.predict(X_test)

# Evaluate: MAE is in units of $100,000 (dataset scale)
mae = mean_absolute_error(y_test, predictions)
r2  = r2_score(y_test, predictions)

print(f"MAE: ${mae * 100_000:,.0f}")
print(f"R-squared: {r2:.3f}")

# Check the learned weights for each feature
for name, coef in zip(data.feature_names, model.coef_):
    print(f"  {name}: {coef:.4f}")
Output
MAE: $51,803
R-squared: 0.576

  MedInc: 0.4487
  HouseAge: 0.0099
  AveRooms: -0.1073
  AveBedrms: 0.6451
  Population: -0.0000
  AveOccup: -0.0038
  Latitude: -0.4214
  Longitude: -0.4333

The model is off by about $52,000 on average and explains 57.6% of the variation in house prices. For a linear model with no feature engineering, that is a reasonable starting point. The coefficients reveal something important: median income (MedInc) has the largest positive weight, meaning higher income areas are the strongest predictor of higher house prices. Latitude and longitude have large negative weights, encoding the fact that houses in coastal California tend to cost more.

You could improve this significantly. Polynomial features, better feature engineering, or switching to a gradient-boosted tree would likely get R-squared above 0.85. But for understanding the core idea, a simple linear model is always the right first step. Understand what a simple model says before reaching for something complex.

"All models are wrong, but some are useful."

George Box, statistician
Hands-on activity

Predict California house prices and explain your results

You will train a linear regression model, inspect what it learned, plot the residuals, and then upgrade to a more powerful model to see how much better it gets. By the end, you will be able to explain in plain English what the model is doing and why its predictions look the way they do.

01 Open the Lesson 3.3 Colab notebook. Run the linear regression pipeline and note the MAE and R-squared.
02 Plot predicted vs actual values using Matplotlib. What would a perfect model look like on this plot? How does yours differ?
03 Plot the residuals (prediction minus actual) against the predicted values. Are the errors random, or do you see a pattern?
04 Replace LinearRegression with RandomForestRegressor(n_estimators=100). Retrain and compare MAE and R-squared. How much did the score improve?
05 Use feature_importances_ from the Random Forest model to see which features matter most. Does the ranking match your intuition about what drives house prices?
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Practice Notebook
Run this lesson's code live in Google Colab
All examples + challenge exercises · Free GPU included · No setup required
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.